In this lesson
Phase 4 · Lesson 4.3

Convolutional Neural Networks

How computers learn to truly see: the elegant architecture behind image recognition, medical imaging, and every face-detection system you use today.

🕑 30 min read 📊 4 visualisations 💻 Keras + MNIST

The Image Problem: Why Dense Networks Struggle

Before we get to the solution, we need to understand the problem. Imagine you have a 224 by 224 pixel colour image (a common size for image models). That image has 224 × 224 × 3 = 150,528 individual pixel values. If you feed these directly into a standard dense (fully connected) layer with just 1,000 neurons, you immediately have 150 million parameters in that one layer alone.

The Parameter Explosion Problem

224 × 224 pixels × 3 colour channels = 150,528 input values
150,528 inputs × 1,000 neurons = 150,528,000 parameters in ONE layer

Training 150 million parameters from a single layer requires enormous amounts of data and compute. And that is only the first layer of many. This approach does not scale.

But there is a second, more fundamental problem. A dense network has no concept of spatial structure. It treats the pixel at the top-left corner and the pixel at the bottom-right corner as completely independent, unrelated inputs. It does not know that nearby pixels tend to be related, that a curve here and an edge there combine to form an eye, or that the same object in two different positions in an image is still the same object.

Dense Network (Fully Connected)

Every input connects to every neuron. No sense of locality. Moving an object slightly to the right looks like a completely different input. 150M+ parameters for one layer.

Convolutional Network (CNN)

Small filters examine local patches of the image. Weights are shared across positions. Far fewer parameters. Naturally handles spatial relationships and translation.

Convolutional Neural Networks (CNNs) were designed specifically to solve these problems. They take inspiration from neuroscience research on the visual cortex, though it is important to note they are a mathematical engineering solution, not a precise biological model.

How Convolution Works

At the heart of a CNN is the convolution operation. The idea is beautifully simple: take a small grid of learnable numbers called a filter (also called a kernel), and slide it across the entire input image. At each position, multiply each filter value by the corresponding pixel value beneath it, sum all the results, and record a single output number. Do this at every possible position, and you get a new grid called a feature map.

Convolution: A Filter Slides Across the Input INPUT (5×5) 1 0 2 3 1 0 4 3 1 0 2 1 5 2 3 1 3 0 4 2 0 2 1 3 1 × FILTER (3×3) 0 1 0 1 1 1 0 1 0 dot product: 0+0+0+4+3+3+0+1+0 = 11 → FEATURE MAP (3×3) 11 ... ... ... ... ... ... ... ... current pos computed next ← Filter slides one step at a time across the image →

The filter (kernel) sits over a 3×3 patch of the input. Each filter value multiplies its matching pixel. All 9 products are summed to produce one number in the feature map. This repeats at every position.

In this example, the filter weights are not set by us manually. They are learned during training by backpropagation, exactly like weights in a regular neural network. The network automatically discovers what filters are useful for the task at hand.

Key Term: Stride and Padding

Stride is how many pixels the filter moves at each step. Stride 1 means it moves one pixel at a time (maximum resolution output). Stride 2 skips every other position, halving the output size. Padding controls whether zeros are added around the border of the input. "Same" padding keeps the output the same size as the input. "Valid" padding (no padding) shrinks the output by the filter size minus one.

Filters and Feature Maps

A single filter produces a single feature map. But in practice, a convolutional layer uses many filters at once. A typical first layer might use 32 filters, each one learning to detect a different low-level feature of the image.

Early Filters Detect

Horizontal edges, vertical edges, diagonal lines, colour gradients, simple textures. The basic visual primitives.

Middle Filters Detect

Corners, curves, circles, specific colour patterns. More complex shapes built from the primitives above.

Later Filters Detect

Eyes, wheels, text patterns, fur textures. High-level semantic features specific to the task.

This hierarchy from simple to complex is one of the most important properties of CNNs. Researchers at OpenAI and Google have visualised the filters learned by large image networks, and the progression from edges to shapes to objects is remarkably consistent across different models trained on different datasets.

How Many Parameters Does a Conv Layer Have?

A convolutional layer with 32 filters, each of size 3×3, on a grayscale (1 channel) input has: 32 filters × 3 × 3 × 1 + 32 biases = 320 parameters, regardless of input image size. Compare this to a dense layer on a 28×28 image output to 32 neurons: 28 × 28 × 32 = 25,088 parameters. The difference grows dramatically with image size.

Pooling Layers: Compressing What We Have Learned

After a convolutional layer produces feature maps, a pooling layer compresses them. The most common type is max pooling: divide the feature map into non-overlapping windows, and keep only the largest value from each window.

Max Pooling (2×2 window, stride 2) FEATURE MAP (4×4) 3 7 1 9 2 4 6 2 5 1 8 3 6 3 4 1 → take max of each window OUTPUT (2×2) 7 9 6 8 max(3,7,2,4) max(1,9,6,2) max(5,1,6,3) max(8,3,4,1)

Max pooling with a 2×2 window and stride 2. Each coloured window in the 4×4 feature map becomes one cell in the 2×2 output. Only the largest value in each window survives.

Pooling achieves three things at once. First, it reduces spatial dimensions (a 4×4 map becomes 2×2), which cuts computation and memory in downstream layers. Second, it provides some robustness to small translations: if a feature is detected slightly to the left or right of a window boundary, the max operation will still capture it. Third, it forces the network to summarise regional information rather than fixating on exact pixel positions.

Max Pooling

Keeps the strongest activation in each window. Good at detecting "was this feature present anywhere in this region?" Most commonly used.

Average Pooling

Averages all activations in each window. Smoother signal, less sharp. Sometimes used in the final layers before classification.

The CNN Architecture: Putting It All Together

A complete CNN stacks convolutional layers, activation functions, and pooling layers into a pipeline that progressively extracts higher-level features, then ends with dense layers that perform the final classification.

INPUT 28×28×1 CONV2D 32 filters 3×3, ReLU 26×26×32 MAXPOOL 2×2 13×13×32 CONV2D 64 filters 3×3, ReLU 11×11×64 MAXPOOL 2×2 5×5×64 FLATTEN 1,600 DENSE 128 ReLU + Dropout OUTPUT 10 Softmax Feature extraction Deeper features Unroll Classify

A CNN for MNIST. Two convolutional blocks extract spatial features, each followed by max pooling to reduce size. The flatten layer unrolls the 5×5×64 feature maps into 1,600 values. The dense layers make the final classification decision.

Notice how the spatial dimensions shrink at each pooling step (28 → 13 → 5) while the number of channels (depth) increases (1 → 32 → 64). This is the classic CNN pattern: compress spatial information while extracting richer feature descriptions. The flatten layer then converts these feature maps into a regular vector that the dense layers can process.

Why CNNs Work: Two Key Principles

1. Local Connectivity

Each neuron in a convolutional layer only connects to a small local patch of the previous layer. A 3×3 filter only "sees" 9 pixels at a time. This makes sense for images: whether a pixel is part of an eye is determined by its immediate neighbours, not by pixels at the opposite corner of the image. Local connectivity dramatically reduces the number of parameters and forces the network to look for local patterns.

2. Weight Sharing

The same filter weights are used at every spatial position in the image. If the network learns a filter that detects a vertical edge, it applies that same detector everywhere, not just at one location. This means the network can recognise a vertical edge whether it appears at the top, bottom, left, or right of an image. This property is called translation equivariance: shifting the input shifts the output by the same amount.

Translation Equivariance vs Invariance

Convolution is technically translation equivariant (shift the input, and the output shifts too). Max pooling then provides some translation invariance within each pooling window (small shifts do not change which value is the maximum). Together, they give CNNs robustness to small positional shifts, which is why a "3" in the top-left corner and a "3" in the bottom-right are recognised as the same digit.

These two properties, combined with the depth of multiple layers, are why CNNs succeeded at image recognition tasks where dense networks previously failed. They encode prior knowledge about the structure of images directly into the architecture, rather than making the network learn everything from scratch.

Landmark CNN Architectures

The history of CNNs is a story of progressively deeper and more creative architectures, each achieving a breakthrough on the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) benchmark, which tests classification across 1,000 categories from 1.2 million training images.

1998
LeNet-5 (LeCun, Bottou, Bengio, Haffner)

Published in Proceedings of the IEEE, this was the first practically successful CNN. It had 5 layers and was trained to read handwritten digits for cheque processing at banks. It established the Conv → Pool → Conv → Pool → Dense pattern that all modern CNNs follow.

2012
AlexNet (Krizhevsky, Sutskever, Hinton)

The inflection point of modern deep learning. AlexNet won the ILSVRC 2012 competition with a top-5 error rate of 15.3%, compared to 26.2% for the runner-up. This gap shocked the computer vision community. AlexNet was the first to train on GPUs and was one of the first to use ReLU activations and Dropout for regularisation at scale.

2014
VGGNet (Simonyan and Zisserman, University of Oxford)

VGGNet showed that depth is a key factor: using only very small 3×3 filters but stacking 16 to 19 layers achieved 7.3% top-5 error. Its clean, uniform design (all 3×3 convolutions, all Max Pool) made it a widely used baseline and teaching tool. The VGG16 and VGG19 models are still used as feature extractors today.

2014
GoogLeNet / Inception (Szegedy et al., Google)

Won ILSVRC 2014 classification with 6.67% top-5 error. Introduced the Inception module, which applies multiple filter sizes (1×1, 3×3, 5×5) in parallel and concatenates their outputs. This allowed the network to capture patterns at different scales simultaneously, all within 22 layers and using far fewer parameters than VGGNet.

2015
ResNet (He, Zhang, Ren, Sun — Microsoft Research)

Published as "Deep Residual Learning for Image Recognition," ResNet introduced skip connections (residual connections) that let gradients bypass layers directly. This solved the vanishing gradient problem and made it possible to train networks with 50, 101, and even 152 layers. ResNet-152 achieved 3.57% top-5 error on ImageNet, surpassing human-level performance on that benchmark (approximately 5%). This architecture remains foundational in 2025.

Surpassing Human Performance on ImageNet

When ResNet beat 5% top-5 error in 2015, headlines declared "AI surpasses human vision." This requires context: ImageNet tests recognising one of 1,000 specific categories, a highly constrained task. Human visual understanding encompasses far more: understanding context, reasoning, reading, spatial navigation, and much else. CNNs also fail in ways no human would, such as being fooled by adversarial examples: images with tiny imperceptible pixel changes that completely fool the model. Performance on one benchmark does not equal general visual intelligence.

Building a CNN in Keras: MNIST Digit Recognition

MNIST is a dataset of 70,000 handwritten digit images (60,000 training, 10,000 test), each 28×28 pixels in grayscale. It has been a standard benchmark since the 1990s. LeCun used it to evaluate LeNet-5, and it remains the "hello world" of CNNs today. A well-tuned CNN typically achieves above 99% accuracy on the test set.

Python (TensorFlow / Keras)
import tensorflow as tf
from tensorflow.keras import layers, callbacks

# ── 1. Load and preprocess MNIST ──────────────────────────────────────
(X_train, y_train), (X_test, y_test) = tf.keras.datasets.mnist.load_data()

# Reshape to (samples, height, width, channels) and scale to [0, 1]
X_train = X_train.reshape(-1, 28, 28, 1).astype("float32") / 255.0
X_test  = X_test.reshape(-1, 28, 28, 1).astype("float32") / 255.0

# ── 2. Build the CNN ──────────────────────────────────────────────────
model = tf.keras.Sequential([
    # Block 1: learn 32 feature maps from 3×3 patches
    # Output: (26, 26, 32) — shrinks by 2 with no padding
    layers.Conv2D(32, (3, 3), activation='relu', input_shape=(28, 28, 1)),
    # Max pool halves spatial size → (13, 13, 32)
    layers.MaxPooling2D((2, 2)),

    # Block 2: learn 64 feature maps from 3×3 patches
    # Output: (11, 11, 64)
    layers.Conv2D(64, (3, 3), activation='relu'),
    # Max pool → (5, 5, 64)
    layers.MaxPooling2D((2, 2)),

    # Flatten the 5×5×64 = 1600 values into a 1D vector
    layers.Flatten(),

    # Dense classification head
    layers.Dense(128, activation='relu'),
    layers.Dropout(0.5),              # drop 50 % of activations during training
    layers.Dense(10, activation='softmax')  # 10 digit classes, probabilities sum to 1
])

# ── 3. Compile ────────────────────────────────────────────────────────
model.compile(
    optimizer='adam',
    loss='sparse_categorical_crossentropy',  # integer labels, not one-hot
    metrics=['accuracy']
)

model.summary()

# ── 4. Train ──────────────────────────────────────────────────────────
early_stop = callbacks.EarlyStopping(
    monitor='val_accuracy', patience=5, restore_best_weights=True
)

history = model.fit(
    X_train, y_train,
    epochs=15,
    batch_size=64,
    validation_split=0.1,
    callbacks=[early_stop]
)

# ── 5. Evaluate ───────────────────────────────────────────────────────
test_loss, test_acc = model.evaluate(X_test, y_test, verbose=0)
print(f"Test accuracy: {test_acc:.4f}")
Model: "sequential" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= conv2d (Conv2D) (None, 26, 26, 32) 320 max_pooling2d (None, 13, 13, 32) 0 conv2d_1 (Conv2D) (None, 11, 11, 64) 18,496 max_pooling2d_1 (None, 5, 5, 64) 0 flatten (None, 1600) 0 dense (Dense) (None, 128) 204,928 dropout (None, 128) 0 dense_1 (Dense) (None, 10) 1,290 ================================================================= Total params: 224,034 ================================================================= Epoch 1/15: loss: 0.2498 - accuracy: 0.9236 - val_accuracy: 0.9840 Epoch 5/15: loss: 0.0651 - accuracy: 0.9800 - val_accuracy: 0.9913 Epoch 10/15: loss: 0.0429 - accuracy: 0.9869 - val_accuracy: 0.9927 ... Test accuracy: 0.9921

Understanding the Output

A few things worth noting from the model summary. The first Conv2D layer has only 320 parameters: 32 filters × (3×3 weights + 1 bias) = 320. Despite seeing the entire 28×28 image, this layer uses far fewer parameters than a dense layer would. The second Conv2D has 18,496 parameters because it processes 32 input channels: 64 × (3×3×32 + 1) = 18,496.

The test accuracy of 99.21% means the network misclassifies only 79 of the 10,000 test images. This is a standard, reproducible result for this class of architecture on MNIST. Note that 224,034 total parameters is quite small by modern standards: large vision models used in production today have hundreds of millions to billions of parameters.

Why sparse_categorical_crossentropy?

There are two cross-entropy losses in Keras for multi-class classification. Use sparse_categorical_crossentropy when your labels are plain integers (0, 1, 2, ..., 9). Use categorical_crossentropy when your labels are one-hot encoded arrays ([1,0,0,...], [0,1,0,...], etc.). MNIST labels from load_data() are integers, so we use sparse.

Try It in Google Colab

This code runs directly in Google Colab without any local installation. Go to colab.research.google.com, create a new notebook, enable GPU (Runtime > Change runtime type > T4 GPU), and paste the code above. Training takes under two minutes on a GPU.

Key Takeaways

Coming Up: Transformers and Attention

CNNs excel at grid-structured data like images. But for sequences (text, speech, time series), a different architecture has become dominant: the Transformer. In Lesson 4.4, you will learn how the self-attention mechanism lets a model weigh the importance of every element relative to every other, enabling ChatGPT, BERT, and modern language models.

Practice Notebook
Run this lesson's code in Google Colab
All examples + challenges · Free GPU included · No setup needed
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.